Back

JMIR Medical Informatics

JMIR Publications Inc.

Preprints posted in the last 30 days, ranked by how well they match JMIR Medical Informatics's content profile, based on 18 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

Lim, H.; Yi, H.; Yoon, J. Y.; Kwon, H.; Lee, D.; Kim, N.

2026-08-21 health informatics 10.64898/2026.08.18.26359085 medRxiv
Top 0.1%
15.0%
Show abstract

Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.

2
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.1%
14.7%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

3
Bridging the "Ten Walls" of Japanese Healthcare Data: A Comprehensive Semantic Mapping of JIPAD to HL7 FHIR R4 and Institutional Gap Analysis for the Japanese Health Data Space (JHDS)

Ohno, K.; Hashimoto, S.

2026-08-10 health informatics 10.64898/2026.08.06.26359847 medRxiv
Top 0.1%
12.6%
Show abstract

Background: Japan faces critical challenges in medical data interoperability, conceptualized as the "Ten Walls" obstructing the Japanese Health Data Space (JHDS) [1]. The Japanese Intensive Care Patient Database (JIPAD) - Japan's largest national ICU registry with 151 participating facilities - represents a high-quality critical care dataset that remains isolated from international data ecosystems. Objective: To develop a formal mapping of all 122 JIPAD variables to HL7 FHIR R4, characterize the nature and magnitude of semantic gaps, and assess the feasibility of JIPAD integration into the JHDS. Methods: All 122 JIPAD variables (Data Dictionary v3.7.2; Linkage Items List 20231020) were evaluated using ISO 21564 [8]-based semantic equivalence scoring across three tiers: High (direct FHIR R4 Core mapping), Partial (mapping via JP-Core Implementation Guide extensions [3]), and Low/No Equivalence (structural institutional gap). Semantically identical multi-instance fields (e.g., secondary disease codes x5) were consolidated into single mapping entries, yielding 114 mapping entries. Pseudonymization architecture was characterized from primary documentation. Results: Of 114 mapping entries representing the 122 JIPAD variables, 97 (85.1%) achieved High Equivalence via LOINC/SNOMED CT, and 12 (10.5%) achieved Partial Equivalence via JP-Core extensions, value-set translation, or FHIR R4 Core extension mechanisms - yielding a combined technical feasibility of 95.6% (109/114). Only 5 entries (4.4%) were classified as Low/No Equivalence, all attributable to Japan's proprietary disease classification system (288 adult codes; 165 pediatric codes) embedded in the DPC reimbursement framework, plus one Japan-specific procedure (PMX endotoxin adsorption) absent from international terminology systems. Variable-level mapping details are provided in Supplementary Table S1. Critically, JIPAD employs pseudonymization with record-linkage capability, enabling 99% DPC data matching - demonstrating that technical and design-level barriers to FHIR integration have already been resolved. Conclusion: JIPAD is technically and architecturally ready for FHIR integration at a 95.6% level. The remaining 4.4% barrier is exclusively institutional - rooted in MHLW policy frameworks governing the DPC disease classification system [6] - rather than technical. FHIR integration would further unlock pharmacoepidemiological and social epidemiological research currently inaccessible due to data isolation. As the sole national ICU registry providing high-acuity anchor data unavailable in general health records, JIPAD integration is essential for a clinically meaningful JHDS by 2027.

4
Adapting Clinical Event Annotation to Dutch Primary Care: An Event Annotation Framework for Post-Acute Infection Syndromes

Mazzucato, S.; Leeuwenberg, A.; van Doorn, S.; van Rosmalen, J.; Slurink, I. A. L.

2026-08-22 health informatics 10.64898/2026.08.19.26360841 medRxiv
Top 0.1%
11.7%
Show abstract

Extracting clinical information from Dutch free-text medical notes requires language-specific annotation resources, yet Dutch primary care lacks a reusable event-annotation framework for infections, post-acute infection syndromes (PAIS), and related symptoms. We adapted the COVID-19 Annotated Clinical Text (CACT) framework to Dutch and applied it to GP notes for PAIS event extraction. The framework has three annotation layers: a DiagnosticExpression typology covering acute infections, post-acute syndromes, and relevant comorbidities; an eleven-subtype Evidence inventory grounded in Dutch primary-care testing practice; and explicit decision rules for the SOEP structure of Dutch general practitioner (GP) notes (Subjective, Objective, Evaluation, Plan), including the distinction between clinician hedging and patient-side hypotheticals. On a 200-note pilot, span-level F1 under the Lybarger criterion reached 0.51 [95% CI: 0.47, 0.55] across six core entities; restricted to spans both annotators noticed, conditional F1 reached 0.78 [0.75, 0.80], indicating that most disagreement stems from annotation coverage rather than label assignment. The adaptation illustrates how an English event-based clinical annotation framework can be extended to a new language and clinical setting, yielding a reusable resource for Dutch clinical NLP; which steps generalise beyond this case (CACT to Dutch primary care) and which are specific to Dutch or PAIS remain to be tested.

5
DBToken: A Database Tokenizer for Medical Event Foundation Models

Shin, I.; McCann, K.; Marino, G.; Siam, U. T.; Li, H.; Stutz, E.; Edara, R.; Loza, A. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360487 medRxiv
Top 0.1%
7.6%
Show abstract

Objectives Transformer models for electronic health records require converting clinical data into token sequences, however standardized tokenization and evaluation frameworks are lacking. We introduce DBToken, an open-source library, and bits-per-row (BPR), a metric for comparing tokenization strategies. Materials and Methods DBToken accepts Medical Event Data Standard (MEDS)-compatible input and supports multiple text, numeric, and temporal tokenization strategies. BPR extends the bits-per-byte metric used in language models to enable comparison across tokenization strategies. Results DBToken efficiently tokenized data across configurations. BPR identified the vocabulary size associated with the best clinical outcome performance and localized differences in numeric tokenization performance by token class. Discussion Optimal tokenization strategies for medical foundation models are a subject of active research. DBToken enables reproducible tokenization experiments, while BPR efficiently screens vocabulary sizes and numeric representations before downstream evaluation. Conclusion DBToken and the BPR metric provide open-source infrastructure for reproducible EHR tokenization and cross-strategy evaluation.

6
Identifying patients with a phenotype consistent with chronic postsurgical pain after hip and knee arthroplasty using robust, scalable k-medoids clustering analysis

Gillam, L.; Doleman, B.; Knaggs, R.; Williams, J.

2026-08-12 orthopedics 10.64898/2026.08.11.26360161 medRxiv
Top 0.1%
7.0%
Show abstract

Background Chronic postsurgical pain (CPSP) affects between 7-23% and 13-44% of patients after hip and knee arthroplasty, respectively. Standardised methods of pain assessment provide superior evaluation of pain, including the Oxford Joint Score Pain Subscale (OJS-PS). We aim to estimate the proportion of patients with a phenotype consistent with CPSP through a k-medoids clustering technique and identify a threshold on the OJS-PS to highlight such patients at a population level. Methods In this cross-sectional study Patient Reported Outcomes Measures data 6-months after hip and knee arthroplasty from 2017 to 2025 were examined. An adapted k-medoid clustering technique utilising subsampling, batch assignment and probabilistic consensus allocated clusters. A receiver operator characteristic analysis identified a threshold on the OJS-PS noting the lowest scoring cluster. Our categorisation was compared to self-reported severe or moderate pain; sensitivity, specificity and accuracy of this categorisation were calculated. Results We analysed 109,542 hip and 113,799 knee arthroplasty patients; three clusters were used in each analysis. After hip arthroplasty: 14.4% of patients were assigned to the cluster with the lowest median OJS-PS of 11 [IQR 8 - 13]. A threshold of 15.5 classified patients as severe or moderate pain with 60.6% sensitivity, 91.0% specificity and 85.7% accuracy. Similarly, after knee arthroplasty, 25.3% were assigned to the cluster with the lowest median OJS-PS of 14 [IQR 11 - 16]. A threshold of 18.5 on the OJS-PS had an 85.4% sensitivity, 88.4% specificity and 87.8% accuracy for classifying patients with self-reported severe or moderate pain. Conclusions This robust and scalable clustering technique on ordinal clinical data estimates the proportion of patients reporting a phenotype consistent with CPSP. On a population level the thresholds identified on the OJS-PS could aid screening for potential CPSP patients 6 months after hip and knee arthroplasties.

7
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation

Gokhale, R.; Kukreja, M.; Kumar, N.; Gourab, K.

2026-08-10 health informatics 10.64898/2026.08.07.26358932 medRxiv
Top 0.1%
5.6%
Show abstract

Background: Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations. Methods: We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs). Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements. Results: CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy. Conclusion: Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.

8
A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.

2026-08-12 bioinformatics 10.64898/2026.08.07.742453 medRxiv
Top 0.1%
5.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

9
CPT/HCPCS Code Recommendation from Clinical Notes: A Comparative Evaluation of AI Methods

Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.

2026-08-31 health informatics 10.64898/2026.08.29.26361731 medRxiv
Top 0.1%
5.5%
Show abstract

Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.

10
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.1%
5.4%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

11
Evaluating Clinical Concept Extraction and Evidence-Bounded Terminology Linking: Multisite Model Comparison and Pilot Ablation Study

Chen, Y.; Popescu, M.

2026-08-24 health informatics 10.64898/2026.08.20.26360740 medRxiv
Top 0.1%
5.4%
Show abstract

Background: Clinical terminology pipelines must first extract candidate spans from narrative notes and then determine whether those spans map to existing concepts or warrant further review. Evaluation is difficult because span boundaries vary between annotators and because downstream decisions depend on the terminology evidence retrieved for each span. Objective: We evaluated clinical concept extraction, terminology linking across controlled evidence conditions, and ontology-extension triage for terms that remained unmatched after initial terminology screening. Methods: We conducted 3 complementary pilot evaluations that used distinct units of analysis and were analyzed separately. Study 1 compared 5 automated extraction pipelines and a union-merge analysis with 2 human annotation sets in 66 deidentified clinical notes from 3 health systems. Agreement was evaluated by exact string matching and BGE-large-en-v1.5 embedding matching. Study 2 evaluated 56 clinical spans, including 28 with reference Unified Medical Language System concepts and 28 adjudicated as unsuitable for ontology extension, under complete retrieval, matched-concept masking, and large language model-only inference, yielding 168 span-condition outputs. The graph retrieval pipeline used BGE-large-en-v1.5 embeddings, and the decision model was Gemma 3 27B. Study 3 applied full vector retrieval to 84 terms previously not matched in either UMLS or BioPortal. Results: In Study 1, interannotator exact-match F1 was 0.29 and embedding-match F1 was 0.75. Automated exact-match F1 scores ranged from 0.07 to 0.17; embedding-match F1 was highest for MedGemma (0.55), followed by Gemma (0.53), sci_md and SciBERT (each 0.43), and Llama 3.3 (0.32). In Study 2, complete retrieval returned a reference-matched link for 28/28 known-concept spans (100%; 95% CI, 87.9%-100%). Masking assigned POSSIBLE_CANDIDATES to all 28; large language model-only inference assigned POSSIBLE_CANDIDATES to 25/28 (89.3%) and LINKED to 3/28 (10.7%). Across the 3 evidence conditions, the same 12/28 unsuitable-extension spans were classified as NOT_MEANINGFUL (42.9%) and the same 16/28 as POSSIBLE_CANDIDATES (57.1%). In Study 3, the pipeline assigned PLAUSIBLE_EXISTING_CONCEPT to all 84 terms, none was flagged for extension, and top-candidate similarity averaged 0.914 (SD 0.027); extension status was not independently adjudicated. Conclusions: Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence. In the follow-up sample, initial nonmatching did not establish ontology novelty: after semantic retrieval, the pipeline classified all 84 terms as plausible existing concepts and proposed none for extension. These findings support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.

12
Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Barrett, L.; Joshi, N.; North, A. S.; Dimitrov, L.; Maughan, E. F.; Ross, T.; Pankhania, R.; Paramjothy, K.; Minty, I.; Farache-Trajano, L.; Smith, S. L.; Mason, K. A.; Bhargava, E. K.; Donnelly, C.; Fatoum, H.; Padiyar, A.; Kader, Z.; Chan, C. H. K.; Schilder, A. G.; Mehta, N.

2026-08-24 otolaryngology 10.64898/2026.08.21.26361031 medRxiv
Top 0.1%
5.2%
Show abstract

Background: Large language models (LLMs) have shown increasing capability in medical knowledge tasks, yet how they perform in extracting structured clinical information from real-world clinical documentation remains uncertain. We evaluated the performance of LLMs relative to medical professionals in extracting SNOMED-coded clinical information from openly available Ear, Nose and Throat (ENT) EHRs from MTSamples, examining both reliability and accuracy metrics. Methods: We evaluated the performance of seven LLMs (including GPT-4o, Claude 3.5, Gemini 1.5 Pro, Gemma 3 and three LLAMA variants) against annotations from fourteen medical professionals who served as both study authors and data annotators. Each annotator independently extracted seven categories of clinical information from 98 publicly available ENT clinical documents: socio-demographics, symptoms, signs, diagnoses, treatments, risk factors, and test results. Standardised medical terminology was enforced through SNOMED-CT code assignment, enabling standardised comparison through Cohen's Kappa. We employed Bayesian hierarchical modelling to test non-inferiority of medic-LLM agreement compared to medic-medic agreement, using Beta distributed likelihood functions with weakly informative priors. Non-inferiority margins of 0.05, 0.10, and 0.15 were assessed with 95% posterior probability thresholds. Results: Cohen's Kappa for inter-rater reliability was 0.752 (95% CI: 0.710 - 0.794) between medical professionals and 0.391 (95% CI: 0.362-0.420) between LLMs and medical professionals. Bayesian analysis showed medic-medic agreement (posterior mean 0.813, 95% CI: 0.755-0.860) exceeded medic-LLM agreement (0.659, 95% CI: 0.633-0.684) by 0.154 (95% CI: 0.091-0.209). Non-inferiority was rejected at all tested margins (delta = 0.05, 0.10, 0.15). Agreement varied by clinical category, with smallest differences for test results and largest for diagnoses. GPT-4o achieved 97.0% precision and 84.9% recall, with a 7.5% false positive rate. Conclusions: Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation. These findings provide evidence-based guidance for LLM deployment in clinical documentation workflows, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

13
Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization

Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.

2026-08-06 health informatics 10.64898/2026.08.04.26359616 medRxiv
Top 0.1%
4.7%
Show abstract

Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.

14
Toward Transportable Acute Kidney Injury Prediction: An Explainable XGBoost Model with Temporal Validation Using MIMIC-IV

Okundaye, D. O.; Isiekwene, C. C.

2026-09-03 health informatics 10.64898/2026.09.01.26360393 medRxiv
Top 0.1%
4.0%
Show abstract

Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.

15
REFINE: Closing the Loop Between Large Language Models and Symbolic Rules in Clinical NLP

Wang, N.; Kakadiaris, A.; Li, C.; Wang, R.; Ahn, J.; Wang, Y.; Fu, S.

2026-08-17 health informatics 10.64898/2026.08.11.26360118 medRxiv
Top 0.1%
3.4%
Show abstract

Symbolic clinical natural language processing (NLP) systems remain widely used for extracting clinical concepts from electronic health record (EHR) narratives, but maintaining rule resources requires extensive manual error analysis and rule refinement. This study investigates whether large language models (LLMs) can assist in identifying extraction errors and generating candidate rules to improve symbolic clinical NLP systems. Using error reports derived from a multi-site evaluation of a previously validated symbolic model for cognitive and neuropsychiatric-related clinical concepts, we developed a human-in-the-loop framework, REFINE. The framework first uses LLMs to classify extraction errors and generate explanatory reasoning, which can then be incorporated into prompts for rule generation. Three LLMs (GPT-5.2, GPT-4o, GPT-4o-mini) were evaluated under four prompting conditions. LLM-generated rule sets improved performance compared with the baseline NLP-CAM system, increasing F1-score from 0.37 to 0.58. These findings suggest that LLMs can support scalable rule refinement for symbolic clinical NLP systems.

16
TabMedQA: From Structured Data to Question-Answer Datasets in Early Clinical Decision-Making

Iturra-Bocaz, G.; Galuscakova, P.; Vedde, S.; Fernandez-Quilez, A.

2026-08-21 urology 10.64898/2026.08.19.26360779 medRxiv
Top 0.1%
3.4%
Show abstract

The rising adoption of Large Language Models (LLMs) and Retrieval Augmented Generation (RAG) in clinical general practice demands datasets that capture realistic early-stage clinical decision-making, where experts must decide on follow-up actions based on sparse, structured patient data. Existing medical Question-Answering (QA) resources primarily address post-diagnostic or specialist settings and rarely reflect how General Practitioners (GPs) document and justify early decisions based on clinical observations from Electronic Health Records (EHRs) and grounded on clinical guidelines. We present TabMedQA, a framework for synthesizing QA collections that emulate how GPs formulate and document decisions in encounter notes during early patient assessments. TabMedQA leverages instruction-tuned LLMs, guided by disease-specific clinical guidelines, to generate full encounter notes composed of a guideline-grounded justification and a corresponding follow-up recommendation directly from structured EHR inputs. The framework further supports RAG-based evaluation, simulating how GPs might consult previous patient encounters to inform new consultations. We demonstrate the application and resulting resource use of TabMedQA on prostate cancer using the publicly available PI-CAI collection and release the resulting PI-CAI QA collection, resource generation templates, and TabMedQA code. To the best of our knowledge, TabMedQA provides the first open framework for creating guideline-grounded, EHR-based QA collections that enable the generation and holistic evaluation of LLM-produced clinical encounter notes, bridging decision-making accuracy with clinical encounter quality in general practice.

17
Python-Streamlit web application to enhance evidence-based medicine education for first year medical students

Patchigolla, V.; Jhand, A. S.; Lee, H. J.; Benjamins, L. J.

2026-08-26 medical education 10.64898/2026.08.23.26361151 medRxiv
Top 0.2%
3.2%
Show abstract

Evidence-based medicine (EBM) concepts are difficult for medical students to grasp. We developed a Python-Streamlit web application providing interactive visualizations to enhance EBM education. Preliminary use with first year medical students demonstrated high engagement and improved conceptual understanding, supporting the feasibility of integrating interactive, web-based tools into EBM curricula.

18
Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data

Yang, Y.; Gu, B.; Hathaway, D. B.; Wyss, R.; Marengo, L.; Gibbons, J. B.; Lyndon, S.; Wu, J.; Chen, Q.; Liu, N.; Wang, P. S.; Celi, L. A.; Bates, D. W.; Lin, J.; Zhou, L.; Yang, J.

2026-08-14 health informatics 10.64898/2026.08.12.26360210 medRxiv
Top 0.2%
3.2%
Show abstract

Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has focused on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to reproduce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Open-source models and reasoning models showed statistically higher stigma rates than proprietary (1.97% vs. 1.60%; p < 0.01) and non-reasoning models (2.35% vs. 1.70%; p < 0.0001), while the stigma rate difference between the general and medical models is not statistically significant (2.00% vs. 1.80%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.304; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of model-task pairs amplifying stigma in the original input notes. Applying prompt engineering as a destigmatizing approach helped reduce model stigma rates by as much as 91.91% without affecting the model performance. This study shows that stigmatizing language generation is common but reducible during LLMs' reasoning traces, suggesting that well-implemented approaches for LLM monitoring and destigmatizing will be essential for healthcare systems to implement.

19
Evaluating Eight Retrieval-Augmented Generation (RAG) Large Language Models' Responses to Clinical Questions: A Comparative Study

Krump, P. A.; Blasingame, M. N.; Koonce, T. Y.; Williams, A. M.; Su, J.; Giuse, N. B.

2026-08-12 health informatics 10.64898/2026.08.10.26360108 medRxiv
Top 0.2%
3.2%
Show abstract

Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.

20
EpiKG2DAG: a Framework for Automated DAG Construction from Biomedical Text

DU, J.; Deng, G.

2026-08-11 health informatics 10.64898/2026.08.09.26360023 medRxiv
Top 0.2%
3.2%
Show abstract

While Directed Acyclic Graphs (DAGs) are essential for causal inference, their construction often relies on expert heuristics, which bypasses systematic evidence synthesis and creates a critical "evidence retrieval gap" in causal modeling. This study introduces EpiKG2DAG, a framework that supports evidence-anchored candidate DAG generation by transforming unstructured biomedical abstracts into structured epidemiological associations. We utilized DeepSeek-V3 to extract exposure-outcome association triplets from 189,266 abstracts and employed SapBERT for semantic normalization against UMLS concepts. The resulting Epidemiological Knowledge Graph (EpiKG) enables the automated identification of candidate confounders, mediators, and colliders based on graph-theoretic motifs and literature-derived evidence. A case study on COVID-19 and AKI demonstrates that the framework uncovers non-obvious confounders, such as air pollution, while ensuring evidence traceability. This work contributes to the field by mitigating the knowledge-acquisition bottleneck and providing a transparent, reproducible foundation for evidence-based causal modeling.